AutoInfraOps: An Agentic DevOps Coordinator for Autonomous Multi-Cloud Monitoring, Optimization, Governance, and Self-Healing Infrastructure Operations
Multi-cloud adoption enables resilience, flexibility and vendor independence, but it also increases operational complexity through heterogeneous interfaces, fragmented moni-toring, inconsistent governance, and difficult incident response. Traditional DevOps and AIOps solutions provide monitoring, automation or optimisation in isolation, but they rarely deliver end-to-end autonomous, governed and explainable multi-cloud infrastructure operations. This paper proposes AutoInfraOps, an agentic DevOps coordinator for autonomous multi-cloud monitoring, optimisation, governance and self-healing infrastruc-ture operations. The framework integrates specialised agents for monitoring, diagnosis, planning, optimisation, governance and remediation. It uses a multi-cloud adapter layer for AWS, Azure and GCP, an event-driven runtime for agent coordina-tion, anomaly detection for SLA violation identification, cost-performance scoring for workload placement, Human-in-the-Loop approval for high-risk actions, and SHA-256 hash-chain audit logging for decision traceability. The system is evaluated using simulated scenarios including AWS latency spike, Azure cost anomaly, GCP availability drop, cross-cloud dependency failure, cascading failure, high-risk failover and rollback events. Experimental results indicate reduced incident response time, improved SLA compliance, cost-aware remediation, and stronger governance visibility. AutoInfraOps demonstrates a practical foundation for autonomous, explainable and policy-aware multi-cloud operations.
Introduction
The paper presents AutoInfraOps, a simulation-safe framework that enhances infrastructure deployment by integrating Large Language Models (LLMs) with DevOps and Infrastructure-as-Code (IaC) workflows. Traditional deployment tools such as Docker, Kubernetes, Terraform, Jenkins, and GitHub Actions automate repetitive tasks but lack contextual reasoning and adaptability, requiring significant manual configuration and troubleshooting.
AutoInfraOps addresses these limitations through a multi-step prompt pipeline that converts natural language deployment requests into structured deployment plans. The framework includes modules for requirement analysis, prompt generation, context enrichment, LLM reasoning, validation, execution simulation, monitoring, deployment history, and feedback refinement. A deterministic validation layer ensures syntax correctness, resource compliance, security, policy adherence, and protection against hallucinated or unsafe LLM outputs before execution.
A full-stack prototype was implemented using React.js, Python REST services, SQLite, and a mock LLM. Instead of executing real infrastructure commands, the system safely simulates deployment processes and generates realistic logs and performance metrics. Experimental evaluation across four deployment scenarios showed that three deployments succeeded while one unsafe deployment was intentionally blocked by the validation module, demonstrating the effectiveness of the governance layer.
The results indicate that AutoInfraOps improves transparency, safety, and deployment planning through intelligent reasoning, monitoring, and continuous feedback. Compared with traditional deployment approaches, the framework offers natural language interaction, adaptive planning, validation-driven governance, simulation-safe execution, and performance-aware feedback, making it a promising solution for secure and intelligent infrastructure automation.
Conclusion
This paper presented AutoInfraOps, an agentic DevOps coordinator for autonomous multi-cloud monitoring, optimi-sation, governance and self-healing infrastructure operations. The framework addresses operational fragmentation in multi-cloud environments by integrating monitoring, diagnosis, plan-ning, optimisation, governance, remediation, verification and auditability into a unified agentic runtime. The proposed system uses a multi-cloud adapter layer for AWS, Azure and GCP, anomaly detection for incident identification, provider scoring for workload placement, HITL approval for high-risk actions and SHA-256 hash chaining for immutable audit records. Experimental evaluation using seven failure scenarios demonstrated reduced MTTR, improved SLA compliance, measurable cost savings and stronger governance coverage compared with manual operations and traditional AIOps. The results indicate that agentic infrastructure operations can provide practical benefits for cloud reliability engineering, DevOps automation and multi-cloud governance. Future work will extend AutoInfraOps using multi-agent reinforcement learning, foundation model agents, autonomous FinOps, edge-cloud coordination, federated agentic operations and digital twin infrastructure environments.
References
[1] P. Notaro, J. Cardoso, and M. Gerndt, “A survey of aiops methods for failure management,” ACM Transactions on Intelligent Systems and Technology, vol. 12, no. 6, pp. 1–45, 2021.
[2] L. Zhang, T. Jia, M. Jia, Y. Wu, A. Liu, Y. Yang, Z. Wu, X. Hu, P. S. Yu, and Y. Li, “A survey of aiops in the era of large language models,” ACM Computing Surveys, vol. 58, no. 2, pp. 1–35, 2025.
[3] M. A. Ali, F. Dornaika, and J. Charafeddine, “Agentic ai: A com-prehensive survey of architectures, applications, and future directions,” Artificial Intelligence Review, vol. 59, no. 1, 2025.
[4] L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin, W. X. Zhao, Z. Wei, and J.-R. Wen, “A survey on large language model based autonomous agents,” Frontiers of Computer Science, vol. 18, no. 6, 2024.
[5] Y. Wang, Y. Pan, Z. Su, Y. Deng, Q. Zhao, L. Du, T. H. Luan, J. Kang, and D. Niyato, “Large model-based agents: State-of-the-art, cooperation paradigms, security and privacy, and future trends,” IEEE Communications Surveys & Tutorials, vol. 28, pp. 1906–1949, 2025.
[6] G. Zhou, W. Tian, R. Buyya, R. Xue, and L. Song, “Deep reinforcement learning-based methods for resource scheduling in cloud computing: A review and future directions,” Artificial Intelligence Review, vol. 57, no. 5, 2024.
[7] M. A. Hady, S. Hu, M. Pratama, Z. Cao, and R. Kowalczyk, “Multi-agent reinforcement learning for resources allocation optimization: A survey,” Artificial Intelligence Review, vol. 58, no. 11, 2025.
[8] O. C. Oyeniran, A. O. Adewusi, A. G. Adeleke, L. A. Akwawa, and C. F. Azubuko, “Ai-driven devops: Leveraging machine learning for automated software deployment and maintenance,” Engineering Science & Technology Journal, vol. 4, no. 6, pp. 728–740, 2023.
[9] S. Y. Long, J. Tan, B. Mao, F. Tang, Y. Li, M. Zhao, and N. Kato, “A survey on intelligent network operations and performance optimization based on large language models,” IEEE Communications Surveys & Tutorials, vol. 27, no. 6, pp. 3915–3949, 2025.
[10] W. Xu, J. Chen, P. Zheng, X. Yi, T. Tian, W. Zhu, Q. Wan, H. Wang, Y. Fan, Q. Su, and X. Shen, “Deploying foundation model powered agent services: A survey,” IEEE Communications Surveys & Tutorials, vol. 28, pp. 1483–1519, 2025.
[11] A. J. Apeh, A. O. Hassan, O. O. Oyewole, O. G. Fakeyede, P. A. Okeleke, and O. R. Adaramodu, “Grc strategies in modern cloud infrastructures: A review of compliance challenges,” Computer Science & IT Research Journal, vol. 4, no. 2, pp. 111–125, 2023.
[12] W. Daniel, G. Andreas, and K. Kilian, “Scaling of end-to-end gover-nance risk assessments for ai systems,” in Dagstuhl Research Online Publication Server, 2025.
[13] Y. Sharma, D. Bhamare, N. Sastry, B. Javadi, and R. Buyya, “Sla management in intent-driven service management systems: A taxonomy and future directions,” ACM Computing Surveys, vol. 55, pp. 1–38, 2023.
[14] A. Tirulo, M. Yadav, M. Lolamo, S. Chauhan, P. Siano, and M. Shafie-Khah, “Beyond automation: Unveiling the potential of agentic intelli-gence,” Renewable and Sustainable Energy Reviews, vol. 226, p. 116218, 2025.